Optimization of a Parallel CFD Code and Its Performance Evaluation on Tianhe-1A
نویسندگان
چکیده
This paper describes performance tuning experiences with a parallel CFD code to enhance its performance and flexibility on large scale parallel computers. The code solves the incompressible Navier-Stokes equations based on the novel Slightly Compressible Model on three-dimensional structure grids. High level loop transformations and argument based code specialization are utilized to optimize its uniprocessor performance. Static arrays are converted into dynamically allocated arrays to improve the flexibility. The grid generator is coupled with the flow solver so that they can exchange grid data in the memory. A detailed performance evaluation is performed. The results show that our uniprocessor optimizations improve the performance of the flow solver for 1.38× to 3.93× on Tianhe-1A supercomputer. In memory grid data exchange optimization speeds up the application startup time by nearly two magnitudes. The optimized code exhibits an excellent parallel scalability running realistic test cases. On 4 096 CPU cores, it achieves a strong scaling parallel efficiency of 77.39 % and a maximum performance of 4.01 Tflops.
منابع مشابه
Performance optimizations for scalable CFD applications on hybrid CPU+MIC heterogeneous computing system with millions of cores
For computational fluid dynamics (CFD) applications with a large number of grid points/cells, parallel computing is a common efficient strategy to reduce the computational time. How to achieve the best performance in the modern supercomputer system, especially with heterogeneous computing resources such as hybrid CPU+GPU, or a CPU + Intel Xeon Phi (MIC) co-processors, is still a great challenge...
متن کاملImproved Algorithm for Reconstructing Singular Connection in Multi-Block CFD Applications
In this paper, an improved algorithm is proposed for the reconstruction of singularity connectivity from the available pairwise connections during preprocessing phase. To evaluate the performance of our algorithm, an in-house CFD code, in which high-order finite-difference method for spatial discretization, running on the Tianhe-1A supercomputer is employed. Test cases with a varied amount of m...
متن کاملPerformance Evaluation of a Ranque-Hilsch Vortex Tube with Optimum Geometrical Dimensions
A brass Vortex Tube (VT) with interchangeable parts is used to determine the optimum cold end orifice diameter, main tube length and diameter. Experim...
متن کاملConstruction and Evaluation of an Incremental Iterative Version of a Parallel Multigrid Cfd Code via Automatic Diierentiation for Shape Optimization Construction and Evaluation of an Incremental Iterative Version of a Parallel Multigrid Cfd Code via Automatic Diierentiation for Shape Optimization
Automatic diierentiation (AD) is a technique for augmenting computer codes to compute derivatives of a subset of their outputs with respect to a subset of their inputs. AD has been shown to provide accurate, but ineecient, sensitivity-enhanced CFD codes for use in aerodynamic shape optimization. To address the ineeciency problem, several special purpose techniques have been suggested. One such ...
متن کاملHeterogeneous Programming and Optimization of Gyrokinetic Toroidal Code and Large-Scale Performance Test on TH-1A
In this work, we discuss the porting to the GPU platform of the latest production version of the Gyrokinetic Torodial Code (GTC), which is a petascale fusion simulation code using particle-in-cell method. New GPU parallel algorithms have been designed for the particle push and shift operations. The GPU version of the GTC code was benchmarked on up to 3072 nodes of the Tianhe-1A supercomputer, w...
متن کاملذخیره در منابع من
با ذخیره ی این منبع در منابع من، دسترسی به آن را برای استفاده های بعدی آسان تر کنید
عنوان ژورنال:
- Computing and Informatics
دوره 33 شماره
صفحات -
تاریخ انتشار 2014